Clustering Gene Expression Profiles Using Mixture Model Ensemble Averaging Approach

نویسنده

  • FAMING LIANG
چکیده

Clustering has been an important tool for extracting underlying gene expression patterns from massive microarray data. However, most of the existing clustering methods cannot automatically separate noise genes, including scattered, singleton and mini-cluster genes, from other genes. Inclusion of noise genes into regular clustering processes can impede identification of gene expression patterns. The model-based clustering method has the potential to automatically separate noise genes from other genes so that major gene expression patterns can be better identified. In this paper, we propose to use the ensemble averaging method to improve the performance of the single model-based clustering method. We also propose a new density estimator for noise genes for Gaussian mixture modeling. Our numerical results indicate that the ensemble averaging method outperforms other clustering methods, such as the quality-based method and the single model-based clustering method, in clustering datasets with noise genes.

برای دانلود رایگان متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

منابع مشابه

Bayesian infinite mixture model based clustering of gene expression profiles

MOTIVATION The biologic significance of results obtained through cluster analyses of gene expression data generated in microarray experiments have been demonstrated in many studies. In this article we focus on the development of a clustering procedure based on the concept of Bayesian model-averaging and a precise statistical model of expression data. RESULTS We developed a clustering procedur...

متن کامل

Mixture model averaging for clustering

Mixture Model Averaging for Clustering Yuhong Wei University of Guelph, 2012 Advisor: Dr. Paul D. McNicholas Model-based clustering is based on a finite mixture of distributions, where each mixture component corresponds to a different group, cluster, subpopulation, or part thereof. Gaussian mixture distributions are most often used. Criteria commonly used in choosing the number of components in...

متن کامل

Clustering of time-course gene expression profiles using normal mixture models with AR(1) random effects

Motivation: Time-course gene expression data such as yeast cell cycle data may be periodically expressed. To cluster such data, currently used Fourier series approximations of periodic gene expressions have been found not to be sufficiently adequate to model the complexity of the time-course data, partly due to their ignoring the dependence between the expression measurements over time and the ...

متن کامل

Bayesian Model-Averaging in Unsupervised Learning From Microarray Data

Unsupervised identification of patterns in microarray data has been a productive approach to uncovering relationships between genes and the biological process in which they are involved. Traditional model-based clustering approaches as well as some recently developed model-based mining approaches for integrating genomic and functional genomic data rely on one’s ability to determine the correct ...

متن کامل

A Mixture model with random-effects components for clustering correlated gene-expression profiles

MOTIVATION The clustering of gene profiles across some experimental conditions of interest contributes significantly to the elucidation of unknown gene function, the validation of gene discoveries and the interpretation of biological processes. However, this clustering problem is not straightforward as the profiles of the genes are not all independently distributed and the expression levels may...

متن کامل

ذخیره در منابع من


  با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید

عنوان ژورنال:

دوره   شماره 

صفحات  -

تاریخ انتشار 2008